Skip to main content

Health Data

"Health data" covers several genuinely different things. Confusing them is the most common cause of dashboards that answer no one's question.


Kinds of health data​

Individual-level clinical data​

Encounters, diagnoses, observations, medications, procedures. Generated during care, structured around the patient. Sensitive by default.

Aggregate routine data​

Counts by facility and period — the classic HMIS model. Cheap to transmit, adequate for coverage monitoring, unable to answer questions about individuals or continuity of care. See DHIS2.

Registry data​

Master lists: clients, facilities, health workers, products. Boring and foundational — most exchange failures are registry failures. See OpenHIE.

Survey and census data​

Population-representative, periodic, expensive. The denominator that routine data usually lacks.

Surveillance data​

Timely, targeted, often incomplete by design; optimised for detecting change rather than measuring level.

Logistics and financial data​

Stock, consumption, claims, expenditure. Frequently the most reliable data in a health system, because money is reconciled.

Patient-generated data​

From apps, wearables and self-report. High volume, variable quality, different consent basis.


Data quality dimensions​

DimensionQuestion
CompletenessDid every reporting unit report?
TimelinessDid it arrive when it was needed?
AccuracyDoes it match the source record?
ConsistencyDo related values agree?
ValidityIs it within plausible ranges and code sets?
UniquenessIs the same person or event counted once?

Quality is produced at the point of collection. Downstream cleaning can detect problems; it cannot create information that was never recorded. The strongest lever is making the data useful to the person entering it.


Denominators​

Routine data supplies numerators. Coverage indicators need denominators — population estimates, target populations, catchment areas — and these are usually the weakest part of any coverage figure. Report the denominator source alongside the indicator, or the number is uninterpretable.


De-identification​

Removing names is not de-identification. Re-identification risk comes from combinations: date of birth, sex, facility and visit date can identify a person in a small population.

Techniques: removing direct identifiers, date shifting, generalising geography and age, aggregating small cells, and formal approaches such as k-anonymity or differential privacy for release.

Small numbers are the recurring hazard in health reporting — a cell of one in a district table can be a disclosure. Suppression rules should be set before publication, not after a complaint.